Papers with up inference
Seq2Edits: Sequence Transduction Using Span-level Edit Operations (2020.emnlp-main)
Copied to clipboard
| Challenge: | Seq2Edits is an open-vocabulary approach to sequence editing for natural language processing tasks with a high degree of overlap between input and output texts. |
| Approach: | They propose an open-vocabulary approach to sequence editing for NLP tasks with a high degree of overlap between input and output texts. |
| Outcome: | The proposed approach speeds up inference by up to 5.2x compared to full sequence models . it improves explainability by associating each edit operation with a human-readable tag. |
Block Pruning For Faster Transformers (2021.emnlp-main)
Copied to clipboard
| Challenge: | Pruning methods have proven to be effective at reducing model size, while distillation methods are proven for speeding up inference. |
| Approach: | They propose a block pruning approach that integrates structured pruning methods with the movement pruning paradigm for fine-tuning. |
| Outcome: | The proposed model is 2.4x faster, 74% smaller and faster than distilled models on classification and generation tasks. |